Evaluation of Background Knowledge for Latent Semantic Indexing Classification

نویسندگان

  • Sarah Zelikovitz
  • Finella Marquez
چکیده

This paper presents work that evaluates background knowledge for use in improving accuracy for text classification using Latent Semantic Indexing (LSI). LSI’s singular value decomposition process can be performed on a combination of training data and background knowledge. Intuitively, the closer the background knowledge is to the classification task, the more helpful it will be in terms of creating a reduced space that will be effective in performing classification. Using a variety of data sets, we evaluate sets of background knowledge in terms of how close they are to training data, and in terms of how much they improve classification.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Sprinkled Latent Semantic Indexing for Text Classification with Background Knowledge

In text classification, one key problem is its inherent dichotomy of polysemy and synonym; the other problem is the insufficient usage of abundant useful, but unlabeled text documents. Targeting on solving these problems, we incorporate a sprinkling Latent Semantic Indexing (LSI) with background knowledge for text classification. The motivation comes from: 1) LSI is a popular technique for info...

متن کامل

Improving Text Classification with LSI Using Background Knowledge

We present work in progress that uses Latent Semantic Indexing (LSI) in conjunction with background knowledge and unlabeled examples to improve text classification accuracy. The singular value decomposition (SVD) that is performed by LSI is done on an expanded term by document matrix that includes the labeled training examples as well as the unlabeled examples. We report classification accuracy...

متن کامل

Integrating Background Knowledge into Nearest-Neighbor Text Classification

This paper describes two different approaches for incorporating background knowledgeinto nearest-neighbor text classification.Our first approachuses backgroundtext to assessthe similarity betweentraining and test documentsrather than assessing their similarity directly. The second method redescribes examples using Latent Semantic Indexing on the background knowledge, assessing document similari...

متن کامل

Modeling and Diagnosing Domain Knowledge Using Latent Semantic Indexing

A Latent Semantic Index (LSI) was constructed from arguments made by Navy officers concerning events in an Anti-Air Warfare scenario. A model based on LSI factor values predicted level of domain expertise with 89% accuracy. The LSI factor space was reduced using MDS to five dimensions: aircraft route, aircraft response, kinematics, localization, and an unclassifiable element. Arguments in the l...

متن کامل

Supervised Locality Preserving Indexing for Text Categorization

A major characteristic of text categorization problems is the prohibitive high dimensionality of the feature space. Most discrimination methods can not work in such a condition, Latent Semantic Indexing (LSI) has been adopted to solve this problem. However, LSI is not an optimal representation for text categorization task mainly because of two reasons: first, the discriminative categorical info...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2005